ci: bump base image/runner to v26.4 and drop v26.3-only workarounds#896
Open
WangLingxun wants to merge 9 commits into
Open
ci: bump base image/runner to v26.4 and drop v26.3-only workarounds#896WangLingxun wants to merge 9 commits into
WangLingxun wants to merge 9 commits into
Conversation
WangLingxun
requested review from
Xiaoming-AMD,
limou102 and
wenxie-amd
as code owners
July 21, 2026 05:57
WangLingxun
marked this pull request as draft
July 21, 2026 06:21
- ci.yaml: BASE_IMAGE + runs-on v26.3 -> v26.4; drop the v26.3-only 'Purge stale TorchTitan' step (v26.4 ships no torchtitan to shadow). Keep 'Install fixed origami' (v26.4 still ships no origami). - Dockerfile: default BASE_IMAGE + comments v26.3 -> v26.4. - docs/scripts: bump default image refs to v26.4, release branch to release/v26.4, and pip pin to primus==26.4.0.
WangLingxun
force-pushed
the
chore/ci-bump-primus-v26.4
branch
2 times, most recently
from
July 21, 2026 09:24
a6e9036 to
90e0389
Compare
- Dockerfile/ci.yaml: drop the origami install; base image ships none, so primus_turbo falls back to heuristic MoE grouped-gemm kernel selection. - aiter_deepbind_patches.py: skip the RTLD_DEEPBIND hook once transformer_engine >= 2.14 (fixed in rocm/primus:v26.4); still applies on older images. - Add unit tests for the new gating logic.
Container reuse compared only the image tag string, not content. On main, IMAGE_TAG is always "latest", so after build-docker repushes under that tag, the check still saw an identical string and reused the stale container instead of pulling the new image. Fix: pull first, then compare docker image inspect .Id (target) against the container's docker inspect .Image (actual) instead of the tag name.
WangLingxun
force-pushed
the
chore/ci-bump-primus-v26.4
branch
from
July 23, 2026 10:49
62f555d to
11e1885
Compare
WangLingxun
marked this pull request as ready for review
July 23, 2026 10:54
v26.4 ships a prebuilt rocSHMEM plus correct ROCM_HOME/UCX_HOME/MPI_HOME for its pip-installed ROCm SDK layout. Rebuilding rocSHMEM from source and re-pinning those vars to their old v26.3 paths made the rebuild CMake step fail (ROCM_HOME pointed at a now-nonexistent /opt/rocm) and is redundant anyway, so drop both. The restore instructions left in the Dockerfile intentionally read the vars from the base image instead of re-declaring them, so they stay correct across future base image bumps.
WangLingxun
force-pushed
the
chore/ci-bump-primus-v26.4
branch
from
July 24, 2026 07:56
7629d86 to
1a48a5b
Compare
Same root cause as the earlier Dockerfile fix, in a file that hadn't been touched yet: ROCM_PATH/MPI_PATH were hardcoded to v26.3's /opt/rocm and system-OpenMPI paths, which no longer exist under v26.4's pip-installed rocm-sdk-devel. rccl/amd-anp also hardcode /opt/rocm internally, and amdclang++ can't find its sibling clang++ (moved under llvm/bin/), so symlink both in rather than patching those repos. Also add pipefail so `cmd | tee log` steps fail on `cmd`'s exit code instead of tee's, which was silently swallowing the above failures until a much later, unrelated- looking step. Verified locally end to end (rccl + amd-anp build install).
Our custom triton build (pinned commit) leaves torch's dist-info Requires-Dist pointing at the base image's stock triton, so later pip installs (e.g. unit-test CI) see it as unsatisfiable and silently replace torch with a plain PyPI/CUDA build. Patch the metadata to the actual triton version right after installing it, and restore the TRITON_COMMIT build-arg that had been dead since #694.
…symbol Primus-Turbo links librocshmem.a via -l: (not --whole-archive) under -fgpu-rdc/--hip-link, so rocSHMEM's team.cpp.o - which defines the plain rocshmem::ROCSHMEM_TEAM_WORLD symbol - never gets pulled into libprimus_turbo_kernels.so. Being a shared object, the missing symbol doesn't fail the build; it only surfaces as an ImportError at runtime. Fix by re-exporting that symbol from a small shim .so and wiring it in via patchelf.
Under PyTorch 2.12's AOTAutograd, letting Dynamo trace into this Function's forward (the default for the modern forward/setup_context split) silently computes wrong gradients for the training-only proxy enabled by tests/unit_tests/.../test_fp8_compile_graph_breaks.py (TestDualFP8Convergence::test_dual_compiled_setup_ctx_vs_eager passed under v26.3/torch 2.10, diverged to rel_diff~0.2-0.3 under v26.4/torch 2.12). Reproduced on real MI300X hardware with backend="eager" (pure Dynamo graph capture, no Inductor codegen) showing the identical divergence, which rules out kernel fusion/codegen and points at Dynamo/AOTAutograd's handling of the 6 forward outputs that exist only to be captured by setup_context/save_for_backward and are otherwise unused by the outer graph. Matches pytorch/pytorch#131794, apparently regressed by the AOTAutograd unused-output backward-pruning rework in pytorch/pytorch#186355. @torch._dynamo.allow_in_graph makes Dynamo treat .apply() as a single opaque node instead of inlining into forward, sidestepping the bug while still keeping everything in one fused graph (verified via torch._dynamo.explain: graph_count=1, graph_break_count=0). Same mechanism the file's own @allow_in_graph baseline classes already use. Verified locally on MI300X: the full tests/unit_tests/backends/megatron/diffusion/test_fp8_compile_graph_breaks.py file (28 tests) now passes with exact eager/compiled parity across eager, aot_eager, and inductor backends, and the full CI unit-test command (pytest --maxfail=1 ./tests/unit_tests/ ... with the existing deselects) passes 1379 passed / 0 failed.
WangLingxun
force-pushed
the
chore/ci-bump-primus-v26.4
branch
from
July 25, 2026 15:22
de44526 to
f9fc67c
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Upgrade the Primus runtime image from
rocm/primus:v26.3tov26.4across CI Dockerfile/workflow, example launch scripts, benchmark scripts and docs, drop v26.3-only CI workarounds now that they're no longer needed, and fix several bugs found while validating the change end-to-end. JAX/MaxText base image not upgraded this round.Image bump
.github/workflows/ci.yaml&docker/Dockerfile:BASE_IMAGE-> v26.4, sync v26.3 commentsrelease/v26.4, pip pin toprimus==26.4.0runs-onstays on thev26.3self-hosted runner label (nov26.4runner registered yet)Workaround cleanup
rocm-libraries@223648a) install from Dockerfile and CI: this base image ships no origami, soprimus_turbofalls back to heuristic kernel selection for MoE grouped-gemm. Left the pip install one-liner in a comment to add it back if ever needed.aiter_deepbind_patches.pyself-detects whether the RTLD_DEEPBIND workaround is still needed: skips installing it oncetransformer_engine.__version__ >= 2.14(fixed in v26.4), still applies on v26.3 and older. Added unit tests for the gating logic.ROCM_HOME/UCX_HOME/MPI_HOMEoverrides: v26.4 already ships a working prebuilt rocSHMEM and correct env vars for its pip-installed ROCm SDK layout; the old overrides pointed at a now-nonexistent/opt/rocmand broke the rebuild.v26.4 pip-installed ROCm SDK fallout
v26.4 switched from a filesystem-installed
/opt/rocmto a pip-installedrocm-sdk-develvenv package. Several things assumed the old layout:Dockerfile.ainic:ROCM_PATH/MPI_PATHwere hardcoded to v26.3 paths; rccl/amd-anp also hardcode/opt/rocminternally, andamdclang++couldn't find its siblingclang++(moved underllvm/bin/). Fixed via symlinks instead of patching those repos, plus addedpipefailsocmd | tee logsteps fail oncmd's exit code (was silently swallowing failures until a much later, unrelated-looking step).Requires-Dist: triton==<exact commit>in its metadata. Our Dockerfile force-reinstalls a different pinned triton commit without updating that metadata, so any laterpip install(e.g. the unit-test job'srequirements.txt) saw torch's declared triton dependency as unsatisfiable and silently replaced torch with a plain PyPI/CUDA build, breaking torchvision and other ROCm extensions. Fixed by patching torch's metadata to the actually-installed triton version right after building it. Also restored theTRITON_COMMITbuild-arg, which had been dead (hardcoded to09500db9in the Dockerfile) since Build Triton from source and bump Primus-Turbo #694.CI fix
docker.io/tasimage/primus:${IMAGE_TAG}), not content. Onmain,IMAGE_TAGis alwayslatest, so afterbuild-dockerrepushes under that tag, the check still saw an identical string and reused the stale container instead of pulling the new image. Fixed by pulling first and comparing the resolved image id instead of the tag name.